Papers with neural audio codecs
UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook (2025.acl-long)
Copied to clipboard
Yidi Jiang, Qian Chen, Shengpeng Ji, Yu Xi, Wen Wang, Chong Zhang, Xianghu Yue, ShiLiang Zhang, Haizhou Li
| Challenge: | Existing neural audio codecs are not capable of handling multi-domain audio data . et al., 2023) integrate speech modality with text-based large language models . |
| Approach: | They propose a unified audio codec with a single codebook to support multi-domain audio data . they propose combining a mix-of-experts strategy and a partitioned domain-adaptive codebook method . |
| Outcome: | The proposed codec outperforms existing codecs on acoustic and semantic representation capabilities. |
Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated significant strides in generating high-quality speech . discretizing speech by neural audio codecs often results in sequences that differ from text sequences . |
| Approach: | They quantitatively analyze the Discrete Representation Inconsistency phenomenon within popular audio tokenizers such as EnCodec. |
| Outcome: | The proposed method mitigates the DRI phenomenon within popular audio tokenizers such as EnCodec. |
Hierarchical Representation Alignment Learning of Diffusion Transformers for Neural Audio Codec (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in diffusion and conditional flow matching models for low-resolution domains are underexplored. |
| Approach: | They propose a CFM-based model that iteratively generates raw waveform in low-bitrate conditions . they propose DVQ, a factorized quantization method that uses a single quantizer . |
| Outcome: | The proposed model outperforms state-of-the-art neural audio codecs in audio quality and semantic intelligibility under low-bitrate conditions. |